文章背景与核心概要
本研究评估了大语言模型(LLM)——特别是 GPT-5.4——在对定性问卷数据进行归纳式内容分析(Inductive Content Analysis)时,相较于人类研究者的表现。研究团队使用了来自欧洲博士生调查中涉及 6 个变量的 903 条开放式回答,并通过调整兰德指数(Adjusted Rand Index, ARI)来评估性能的一致性。
研究发现,LLM 展现出了与人类高度吻合的表现,其编码和主题生成的 ARI 值分别达到 0.61 和 0.54。人类与 LLM 之间的契合度,非常接近人类内部的一致性水平(\(\text{ARI} = 0.68\))以及 LLM 自身的重复一致性水平(\(\text{ARI} = 0.76\))。此外,分析表明模型与人类的一致性高度依赖于数据本身的特性:内部一致性较低的变量,其人机间的契合度也普遍较低。总体而言,该研究证明了大语言模型作为定性研究中可扩展支持工具的巨大潜力。
LLMs for Survey Text Analysis: A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis
LLMs for Survey Text Analysis: A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis
📋 Summary
📋 Summary
本研究评估了大语言模型(LLMs)——具体为 GPT-5.4——在对定性问卷数据进行归纳式内容分析时,相较于人类研究者的能力表现。该研究利用了欧洲博士生调查中 6 个变量的 903 条开放式回答,通过调整兰德指数(ARI)评估了性能对齐情况。
This study evaluates the capability of Large Language Models (LLMs)—specifically GPT-5.4—to conduct inductive content analysis on qualitative survey data compared to human researchers. Using 903 open-ended responses across six variables from a European PhD student survey, the research assessed performance alignment through the Adjusted Rand Index (ARI).
核心要点包括: * 高度对齐: LLM 能够高度贴近人类的表现,在编码和主题生成上分别实现了 0.61 和 0.54 的 ARI 值。 * 内部一致性: 人类与 LLM 之间的契合度,非常接近人类内部的一致性(\(\text{ARI} = 0.68\))以及 LLM 自身的稳定性(\(\text{ARI} = 0.76\))。 * 数据依赖型可靠性: 一致性会随具体变量的不同而大幅波动;内部一致性较低的实体,其人机间的一致性也持续偏低,凸显了数据特征对性能的驱动作用。 * 可扩展性: LLM 展现出强大的潜力,可作为处理归纳式编码任务的定性研究人员的可扩展支持工具。
Key takeaways include: * High Alignment: The LLM closely approximated human performance, achieving ARI values of 0.61 for coding and 0.54 for theme generation. * Internal Consistency: Human-to-LLM agreement fell closely in line with internal consistency rates among humans (\(\text{ARI} = 0.68\)) and within the LLM itself (\(\text{ARI} = 0.76\)). * Data-Dependent Reliability: Agreement fluctuated widely depending on specific variables; entities with low internal consistency consistently showed low between-entity agreement, highlighting how data characteristics drive performance. * Scalability: LLMs demonstrate strong potential as scalable support tools for qualitative researchers managing inductive coding tasks.
📌 Document Metadata
📌 Document Metadata
- arXiv ID:
arXiv:2608.22417[cs.AI] - 主题分类: 人工智能(
cs.AI);计算与语言(cs.CL);人机交互(cs.HC) - 提交日期: 2026年8月23日
- DOI: 10.48550/arXiv.2608.22417
- arXiv ID:
arXiv:2608.22417[cs.AI]- Subjects: Artificial Intelligence (
cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)- Submission Date: 23 August 2026
- DOI: 10.48550/arXiv.2608.22417
作者
- Leonardo Bergmann
- Renata Gheorghiu
- Ana Gvritishvili
- Alex Mican
- Chris Stewart
- Topias Tolonen-Weckström
Authors
- Leonardo Bergmann
- Renata Gheorghiu
- Ana Gvritishvili
- Alex Mican
- Chris Stewart
- Topias Tolonen-Weckström
📄 Abstract
📄 Abstract
大语言模型(LLMs)正越来越多地被用于支持定性研究中的文本分析,然而关于其在归纳式内容分析中表现的实证研究仍然有限。本研究对人类与基于 LLM 的归纳式编码进行了比较,分析对象来自欧洲博士生调查中 6 个变量的 903 条开放式问卷回答。
Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey.
5 名人类编码员遵循标准化的编码方案进行了归纳式内容分析,而一个大语言模型(GPT-5.4)则通过既定的提示词流程执行了相同的任务。使用调整兰德指数(ARI)评估了人类与 LLM 输出之间的一致性。结果显示,人类与 LLM 之间高度对齐,编码的 ARI 值为 0.61,主题生成的 ARI 值为 0.54。这些数值接近人类内部编码和主题结果的一致性(\(\text{ARI} = 0.68\))以及 LLM 自身的一致性(\(\text{ARI} = 0.76\))。
Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (\(\text{ARI} = 0.68\)) and the LLM (\(\text{ARI} = 0.76\)).
不同变量之间的一致性差异较大,实体内部一致性低 consistently(持续)与实体间一致性低相联系,这凸显了数据特征和个人表现对可靠性的作用。总体而言,研究结果表明,在这种特定案例的环境下,LLM 可以近似人类的编码工作(特别是在编码层面),并有望成为归纳式定性分析的可扩展支持工具。
Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.
🔗 Access & Resources
🔗 Access & Resources
- 全文 PDF: 查看 PDF
- 外部文献工具:
- NASA ADS
- Google Scholar
- Semantic Scholar
- Full-Text PDF: View PDF
- External Bibliographic Tools:
- NASA ADS
- Google Scholar
- Semantic Scholar